Visit StickyLock

AI framework helps machines understand 3D spaces better

AI framework helps machines understand 3D spaces better
AI Framework Helps Machines Understand 3D Spaces

For a service robot moving through an office, recognising a group of objects as furniture may provide only a basic understanding of its surroundings. To plan a route or interact with objects, the machine may also need to distinguish a chair from a table or a bookcase, and identify whether an opening is a door or a window. Such distinctions can affect how a machine moves, avoids obstacles and determines which objects it can safely interact with.

Researchers led by Assistant Professor Na Zhao, principal investigator at the Singapore University of Technology and Design (SUTD), and collaborators from Southwest Jiaotong University have developed an artificial intelligence (AI) framework designed to help machines interpret three-dimensional (3D) scenes at multiple levels of detail.

Known as ML3DHS, the framework focuses on 3D hierarchical semantic segmentation. The method assigns several related labels to every point in a 3D scene, allowing a point to carry a broad category such as furniture alongside a more specific classification such as chair, table or sofa.

The work was presented at the 43rd International Conference on Machine Learning in 2026, titled “Multi-Label Learning with Contrastive Cluster Self-Supervision for 3D Hierarchical Semantic Segmentation”.

Three-dimensional vision systems help machines interpret spaces captured by sensors such as LiDAR (Light Detection and Ranging) and depth cameras. The technology is used in embodied intelligence, including robots that navigate buildings, autonomous vehicles that interpret road scenes, and augmented reality applications that overlay digital information onto physical spaces.

Many existing 3D segmentation models assign a single label to each point in a scene. This flat approach can limit how machines represent objects in real-world environments, where the same area may need to be understood at different levels of detail. A model may first identify a set of points as furniture and then determine whether those points belong to a chair, table, sofa or bookcase.

The researchers identified a training challenge because predictions at different levels can compete. A broad prediction may be easier to learn, while a finer prediction requires the model to recognise more detailed distinctions. When these levels share too many model components, improving one prediction can affect how well the other level learns.

ML3DHS addresses this by letting the model share basic 3D information, including shape and geometry, while using separate components for different levels of detail. The framework also uses broad predictions to guide finer ones. Recognising a region as furniture can therefore help the model determine whether it represents a chair, table or sofa.

The system also checks that broad and fine labels remain consistent. This allows the framework to treat the related classifications for each 3D point as a multi-label learning problem, rather than forcing the model to represent only one category at a time.

Another challenge in 3D scene understanding is class imbalance. Common classes such as walls and floors can contain far more points than smaller or less common objects, including windows, doors, columns and clutter. This imbalance can cause models to focus heavily on dominant classes and miss rarer objects that may still matter for navigation and interaction.

ML3DHS addresses this issue through an additional training component that helps the model distinguish object classes more clearly. The component provides a stronger learning signal for rarer objects, supporting the model’s ability to recognise classes that occur less frequently in a scene.

The team evaluated ML3DHS on established indoor and outdoor 3D scene benchmarks and tested it with several AI model architectures. Across these settings, the framework consistently outperformed previous approaches.

On one indoor benchmark, ML3DHS improved segmentation accuracy by 3.38 percentage points compared with the best previous method. The gains were especially notable for less common objects, including clutter, windows, doors and columns.

For an indoor service robot, the distinction between broad and fine categories can provide different levels of environmental information. A broad classification can identify furniture or an opening. In contrast, a finer classification can distinguish a door from a window, a chair from a table, or a column from movable clutter. These distinctions are relevant when a machine needs to choose a safe path or decide which objects it can interact with.

The current framework also has a limitation related to model size during training. ML3DHS increases model size because it uses non-shared decoders and an additional auxiliary branch.

Future work could explore lighter, more efficient models while improving recognition of extremely rare classes. The researchers identified these areas as open questions for further development.

ML3DHS provides a structured approach to recognising objects across different levels of detail in 3D scenes. By combining shared information about shape and geometry with separate processing for different classification levels, the framework addresses both hierarchical label relationships and class imbalance. The work could contribute to more reliable machine perception for robots, autonomous systems and interactive 3D technologies.

Join the Discussion


Visit StickyLock
Back to top